Papers with Google Image Search
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing vision-language models struggle with reasoning-focused tasks due to the lack of high-quality training data. |
| Approach: | They propose a new approach that leverages search engines to create a multimodal multimodal dataset . they use a set of 30,000 seed images to extract HTML data from 700K unique URLs . |
| Outcome: | The proposed model achieves the best known performance on MMMU-Pro (40.7), MathVerse (42.6), and DynaMath (55.7). |
Evaluating the WordsEye Text-to-Scene System: Imaginative and Realistic Sentences (L18-1)
Copied to clipboard
| Challenge: | WordsEye is a system for automatically converting natural language text into 3D scenes representing the meaning of that text. |
| Approach: | They evaluate WordsEye's output vs. simple search methods to find a picture to illustrate a sentence. |
| Outcome: | The WordsEye system produced imaginative sentences and realistic sentences . the results show that wordsEyed produced better results than standard image search engines . |